Papers with large-scale multimodal pre-training

2 papers
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing fine-tuning and continual learning methods compress visual representations and emphasize task alignment over visual retention.
Approach: They propose a modality-decoupled gradient descent (MDGD) that regulates gradient updates to preserve effective rank of visual features and explicitly disentangles visual learning from task-specific alignment.
Outcome: The proposed model reduces visual forgetting and improves visual retention . it disentangles visual learning from task-specific alignment and preserves effective rank .
Modular and Parameter-Efficient Multimodal Fusion with Prompting (2022.findings-acl)

Copied to clipboard

Challenge: Recent research has made impressive progress in large-scale multimodal pre-training.
Approach: They propose to use prompt vectors to align multimodal modalities by pretraining text inputs with prompts or embedding vectors.
Outcome: The proposed method achieves comparable performance to several other multimodal fusion methods in low-resource settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations